Papers by Sean Timothy Okonsky

1 papers
PEaCE: A Chemistry-Oriented Dataset for Optical Character Recognition on Scientific Documents (2024.lrec-main)

Copied to clipboard

Challenge: Existing open-source OCR models focus on scientific texts or generic printed English . Nougat is unable to parse tables in PubMed articles .
Approach: They propose to train OCR models for scientific or generic printed English . Nougat is a popular tool for parsing academic documents, but unable to parse PubMed tables .
Outcome: The proposed models perform better when trained on real-world records than those trained on synthetic records.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations